Papers with typo correction
Building a Japanese Typo Dataset from Wikipedia’s Revision History (2020.acl-srw)
Copied to clipboard
| Challenge: | Typographical errors (typos) also occur in user generated content (UGC). |
| Approach: | They extract over half a million Japanese typo–correction pairs from Wikipedia’s revision history and combine character-based extraction rules, morphological analyzers to guess readings, and various filtering methods to address these challenges. |
| Outcome: | The proposed dataset extracts over half a million typo–correction pairs from Wikipedia’s revision history. |